adapter: remove ENABLE_CLUSTER_CONTROLLER and the legacy REFRESH scheduler - #38101
Conversation
…duler The cluster controller has been default-on since v26.29 and owns the replica set of every user managed cluster. The break-glass dyncfg kept two legacy paths reachable: the REFRESH scheduler in `cluster_scheduling.rs`, whose two entry points already returned before doing any work while the gate was on, and the legacy branches in the ALTER sequencer. Delete the gate and everything only it kept alive: the scheduler module, its coordinator plumbing (two messages, the timer, the select-loop tick, the `cluster_scheduling_decisions` state), the two scheduler metrics, the `cluster_check_scheduling_policies_interval` system var, and the `ReplicaCreateDropReason::ClusterScheduling` variant. The persisted audit vocabulary stays: `SchedulingDecisionsWithReasonsV2` and friends are written by the controller's on-refresh path too, and old events must remain decodable.
f283eea to
ed282e5
Compare
| # period. A zero value would therefore crash-loop environmentd on every boot, | ||
| # so it must be rejected at ALTER SYSTEM SET time. | ||
| simple conn=mz_system,user=mz_system | ||
| ALTER SYSTEM SET cluster_check_scheduling_policies_interval = '0s'; |
There was a problem hiding this comment.
hmm, does our new controller actually run on a similar interval to where we had this before? Now I'm curious.
|
Close, but slower, and the difference is absorbed by a knob that was always the real one. The old poll was The controller's interval is only a fallback cadence, Why the 2s doesn't worry me: the window decision is so the turn-on already leads the refresh by Worth saying explicitly: both numbers are the polling interval, not a bound on how long a refresh takes. Turning the cluster on is the cheap part. Happy to drop |
mtabebe
left a comment
There was a problem hiding this comment.
Change looks good for me. Thanks for updating all the comments too
|
tyty! 🙇♂️ you probably saw, and I don't want to hurry you, but there are some more cleanups left in this stack 😅 |
…nc#38102) Part 2 of 3 of the [design][design] for removing the legacy cluster paths. Stacked on MaterializeInc#38101. Part 3 stacks on this one. [design]: https://github.com/MaterializeInc/materialize/blob/aljoscha/cluster-legacy-05-design/doc/developer/design/20260724_cluster_controller_legacy_removal.md ## Why System/builtin clusters were excluded from the controller's ownership (`ManagedClusterIds` and the sequencer's two ownership tests all required `is_user()`), so the sequencer had to keep a second, complete replica-materialization implementation alive just for them. The exception was originally load-bearing: the boot-time builtin replica migration and the controller would have been two conflicting writers of one replica set. MaterializeInc#37929 removed that conflict. `reconcile_builtin_cluster_replicas` now converges a builtin cluster's replica set on the cluster's own managed config, the same config the controller derives its targets from, so the two converge on the same set by construction. ## What lands Drop the three `is_user()` conjuncts. Runtime ALTERs of system clusters now flow through the controller like any other managed cluster: a config-shape change reshapes into a durable reconfiguration record, a factor change updates the config and the controller converges the replica set within a tick. Ownership and the boot migration compose in both directions. The controller matches replicas by shape and count, never by name, so the migration-created `r1..rN` satisfy its baseline and a steady system cluster reconciles to no decisions. The migration converges by canonical name, so a boot after a reshape (or with a reconfiguration in flight) renames or re-creates replicas the controller materialized under generator names. That is harmless churn: every replica is a cold process at boot anyway, so the cost is replica-id and audit-log noise, and an in-flight record is durable, so the controller picks the reconfiguration back up on its first tick. Boot ordering and 0dt read-only behavior are unchanged. `reconcile_builtin_cluster_replicas` still runs at catalog open, which is what gets system clusters up and hydrated before the serve loop starts and before a 0dt cutover. The controller cannot cover either window: `Coordinator::bootstrap` runs before the controller task is spawned, and the controller is inactive while read-only. Also rewrites the comment in `validate_reconfiguration_resource_limits`. The early return for non-user clusters stays, since system clusters are exempt from `max_replicas_per_cluster` and credit accounting everywhere else, but the reason it gave ("a system cluster never reshapes into a record") is now false. ## Behavior changes - `ALTER CLUSTER mz_system SET (SIZE ...)` goes from a synchronous whole-set recreation to a background graceful reconfiguration (default deadline 24h, `ROLLBACK` on timeout). Strictly more capable: the graceful path did not exist for system clusters before, and part 3 adds `WITH (WAIT FOR '0s')` for operators who cannot afford the overlap set. - `ALTER CLUSTER mz_support SET (REPLICATION FACTOR 1)` updates the config and returns; the controller materializes the replica within a tick. Same as user clusters today. ## Tests - `test/sqllogictest/system-cluster.slt`: the block asserting a replica count right after a factor flip is reworked. slt does not retry, so it now asserts the declared factor (synchronous) and phrases the system-id property as "no replica that isn't system-id'd", which holds regardless of convergence. - `test/testdrive/system-cluster.td` (which does retry) picks up the convergence coverage: factor up materializes exactly one system-id replica, a `SIZE` reshape cuts the realized config over and leaves exactly one replica at the new size, and factor back to 0 retires it with no activity left behind. - `src/environmentd/tests/bootstrap_builtin_clusters.rs`: the three assertions that immediately follow a runtime `ALTER CLUSTER` now poll (`await_declared_and_actual`). The bootstrap and post-restart assertions stay exact, since the catalog-open migration converges synchronously. - `test/cluster/resources/resource-limits.td` already uses retrying queries for its `mz_analytics` factor flips, so it survives unchanged and pins the system-cluster limit exemption end to end. - Docs: the self-managed troubleshooting guide's `mz_catalog_server` resize walkthrough gets a note that the resize is a graceful reconfiguration and `SHOW CLUSTERS` settles rather than flips.
Part 1 of 3 of the design for removing the legacy cluster paths.
Stacked on #38078 (the design doc). Parts 2 and 3 stack on this one.
Why
The cluster controller has been default-on since v26.29 and owns the replica
set of every user managed cluster in production. It still lives behind the
break-glass dyncfg
ENABLE_CLUSTER_CONTROLLER, and that flag is the only thingkeeping the legacy REFRESH scheduler reachable: both of its entry points return
before doing any work while the gate is on, so it is a strict no-op in
production today.
Removing the gate means reverting to the legacy paths requires a binary
rollback. The gate has several releases of burn-in behind it, and the direct
reshape path (see part 3) remains as the operational escape hatch.
What lands
ENABLE_CLUSTER_CONTROLLERdyncfg: definition, registration, its fiveread sites, and the sqllogictest binary's force-on entry.
src/adapter/src/coord/cluster_scheduling.rswholesale:check_scheduling_policies,check_refresh_policy,handle_scheduling_decisions, and theSchedulingDecision/RefreshDecisiontypes.Message::CheckSchedulingPoliciesandMessage::SchedulingDecisionswith their handlers, the timer, theselect-loop tick, and the
cluster_scheduling_decisionsstate.mz_check_scheduling_policies_secondsandmz_handle_scheduling_decisions_seconds(and their rows in the generateddoc/user/data/metrics.yml).cluster_check_scheduling_policies_interval, whose soleconsumer was the timer.
ReplicaCreateDropReason::ClusterScheduling, constructed only by thescheduler.
The persisted audit vocabulary stays:
SchedulingDecisionsWithReasonsV2andfriends are written by the controller's on-refresh path too
(
refresh_window_decision_to_audit_log), and old events must remain decodableregardless.
Tests
test/sqllogictest/mz_cluster_schedules.slt: thecluster_check_scheduling_policies_intervalvalidation block goes with thevar.
test/testdrive/cluster-controller.td: thecc_handoff*legacy-handoffscenarios and the
cc_strandbreak-glass section are deleted (both aregate-off scenarios). The freeze-the-controller technique in the
unmanaged-conversion refusal block is replaced by cranking
cluster_controller_tick_intervalup. Part 3 adds a replacement forcc_strandthat exercises the direct cut-over escape hatch instead.test/pg-cdc/cluster-graceful-reconfiguration.td: the explicit legacyforeground section goes, the controller section stays.
misc/python/materialize/mzcompose/__init__.pypins the flag only forversions below v26.38, so mixed-version runs against older binaries still
exercise the same path.
parallel_workload/action.pydrops it from thedo-not-flip list, and the two launchdarkly-flag-consistency allowlist entries
go.